Generate graph-driven ORT GenAI decoder configs - #611
Conversation
Performance Comparison
|
d673f9d to
62583ae
Compare
Derive generic decoder inputs, outputs, and cache topology from the optimized ONNX graph instead of architecture-name registration. Preserve runtime-specific behavior, fail closed on unsupported state layouts, and publish released-version compatibility metadata.\n\nValidate the exact SmolLM GGUF route with deterministic generation on ORT GenAI 0.14.1 and 0.15.2.\n\nCo-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Limit recorded runtime coverage to onnxruntime-genai 0.15.2 and keep the real SmolLM generation test independent of the ONNX GenAI runtime-evidence packaging path. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
c684c1e to
f43a31f
Compare
There was a problem hiding this comment.
Pull request overview
This PR updates Mobius’s ORT GenAI export path to emit the released, architecture-neutral model.type: "decoder" contract for compatible single-graph decoder-only packages by deriving decoder I/O (including KV-cache templates and cache-slot indexing) directly from the optimized ONNX graph. It also records runtime compatibility metadata and adds end-to-end validation that the evidenced SmolLM GGUF route can load and deterministically generate via ORT GenAI’s generic decoder.
Changes:
- Add graph-driven generic-decoder ABI inspection to derive semantic inputs, outputs, cache name templates, and global cache-slot count; fail closed for unsupported state topologies and incompatible GPT-2 cache ABI.
- Extend
GenaiConfigGeneratorto accept explicit decoder outputs and optionalsliding_windowmetadata, and emitruntime_compatibility.json. - Add/adjust unit + integration tests for the generic decoder contract and SmolLM ORT GenAI generation; thread
runtime_versionthrough GGUF runtime packaging.
Reviewed changes
Copilot reviewed 6 out of 6 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
| tests/gguf_small_model_runtime_integration_test.py | Adds an ORT GenAI integration generation test for the evidenced SmolLM GGUF route using the generic decoder type. |
| src/mobius/integrations/ort_genai/genai_config.py | Allows explicit decoder outputs and optional sliding-window metadata in generated genai_config.json. |
| src/mobius/integrations/ort_genai/auto_export.py | Implements generic-decoder type selection, decoder ABI introspection from the ONNX graph, sliding-window derivation, and writes runtime_compatibility.json. |
| src/mobius/integrations/ort_genai/auto_export_test.py | Adds/updates tests validating the generic decoder ABI, schema invariants, and runtime compatibility metadata emission. |
| src/mobius/integrations/gguf/_runtime_package.py | Threads runtime_version into ORT GenAI config generation during GGUF runtime packaging. |
| docs/cli_reference.md | Documents the generic decoder contract, exceptions requiring specialized runtime types, and emitted runtime_compatibility.json. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| write_ort_genai_config( | ||
| captured[0], | ||
| str(output_dir), | ||
| hf_model_id=case.tokenizer_repository, | ||
| revision=case.tokenizer_revision, | ||
| runtime_version=version("onnxruntime-genai"), | ||
| ) |
Remove the legacy dual-version matrix and validate only the latest stable 0.15.2 release. Refresh the SmolLM route and runtime package evidence after the #611 squash merge, and keep the real artifact lane limited to scheduled, manual, or explicitly requested runs. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
## Summary - add a network-free CPU E2E matrix against pinned `onnxruntime-genai` 0.14.1 and 0.15.2 releases - exercise tokenizer loading, generic decoder model creation, prefill, cache-dependent multi-step decode, deterministic output, reload, and malformed-config rejection - add a scheduled/manual real SmolLM-135M F16 CPU lane using immutable GGUF/tokenizer revisions, verified sizes/hashes, isolated HF/Xet caches, and the production `--runtime ort-genai` package path - enroll runtime routes through schema-validated YAML evidence and expand affected-model selection across graph, task, tokenizer, qtype, runtime helper, workflow, and dependency surfaces - document tiering, exact commands, CPU scope, released capability differences, and the CUDA waiver ## Local validation - OGA 0.14.1 fast lane: 2 passed - OGA 0.15.2 fast lane: 2 passed - OGA 0.15.2 pinned SmolLM F16 real lane: 1 passed - schema/evidence/affected detection: 340 passed - broad non-integration suite: 7,106 passed, 52 skipped, 1 subtest passed - `lintrunner f --output oneline --all-files && lintrunner -a`: clean - two independent security/reliability reviews plus a final post-fix review completed; all actionable findings resolved ## Runtime scope Required CI is CPU-only. CUDA remains scheduled/manual-only because the existing GPU infrastructure uses a CUDA-12-specific prerelease ORT feed and does not provide a stable pinned released OGA lane; this PR makes no CUDA EP claim. ## Stack Based exactly on `justinchuby-generic-genai-configs` / #611. --------- Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com> Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
## Summary - add a dedicated Kimi-K3 config, model, and heterogeneous-state task with exact KDA/NoPE gated-MLA scheduling, AttnRes mixing, SiTU latent MoE, shared experts, and an untied output head - import pinned llama.cpp `kimi-k3` GGUF metadata and tensors with strict metadata/tensor/shape/storage closure and malformed-input rejection - preserve lossless quantization routes under #609: separate rank-3 MLA projections fail closed unless explicitly dequantized, while fused Q4_0 KV-B is split by exact packed-row reordering across weights, scales, and zero-points - preserve valid KDA convolution history across padding, validate kernel/state/task contracts, and cover replay, reorder, float/quantized import, roundtrip, CLI, and negative paths - keep generic OGA runtime packaging truthfully deferred under #605 because released cache schemas cannot represent the heterogeneous state ABI ## Reconstruction Reconstructed after #619 was admin squash-merged. The PR contains only the four-commit Kimi-K3 delta and review follow-ups on live-main base `04b7e3f6f2d9de5d5741aafb8fa6375a18eee693`. This preserves #611 graph-driven ORT GenAI decoder configs, #618 ORT GenAI end-to-end CI, #624 generic decoder config migration, and the Kimi Linear, MiniMax, runtime, and fail-closed quantization changes already on main. `git range-diff` reports all four replayed commits as patch-identical (`=`). ## Validation - Kimi-K3 model and GGUF tests: 30 passed - focused builder/build-graph Kimi-K3 checks: 4 passed, 1900 deselected - affected generic ORT config/runtime/E2E tests: 250 passed - affected OGA metadata tests: 94 passed, 1 skipped - broad non-integration suite: 7765 passed, 56 skipped, 1 subtest passed - repository lintrunner passed - final GPT-5.6 Sol medium review reported no findings ## Real-checkpoint note The pinned `yujiepan/kimi-k3-tiny-random@a5c86ee03f07f7b141508b0108304a1447fbb345` checkpoint was assessed, but macOS cannot run its required `fla-core` kernels and its selective compressed-tensors MXFP4 expert representation is unsupported by the generic HF loader. That format fails explicitly rather than producing an incorrectly quantized graph. --------- Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com> Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
## Summary - clean up still-valid typing, documentation, test robustness, error-message, maintainability, and performance findings left by the GGUF PR stack - preserve the behavioral fixes merged in #625-#632 - hash reused multi-GB GGUF sources once, at the final pre-publication integrity gate, while retaining cheap identity checks around staging Exact base: `2db9d33debdc254d879a51b14434c9a81c230f4f` Exact head: `4541ed2bc9d2ab4227484510b4a85b2d9113eb25` ## Reconstructed original 17-item low-priority tranche The persisted audit retained only the totals, so this list was reconstructed from the live threads and current source. All but the already-fixed #596 comment are addressed in this PR. | PR | Comment | Disposition | |---|---:|---| | #550 | 3837144656 | Implemented: correct tuple return annotation | | #552 | 3837389183 | Implemented: multichannel waveform shape docs | | #559 | 3837716261 | Implemented: name-based cache assertions | | #573 | 3854308350 | Implemented: one final GGUF hash, with integrity regression coverage | | #574 | 3854345521 | Implemented: `TensorRole \| None` typing | | #574 | 3854345597 | Implemented: removed obsolete verdict filtering | | #577 | 3854472859 | Implemented: documented SSM sequence length | | #578 | 3843112613 | Implemented: documented F64 passthrough | | #579 | 3854518332 | Implemented: generalized fused-projection error | | #580 | 3854576300 | Implemented: removed brittle node counts | | #583 | 3854681386 | Implemented: documented conditional draft outputs | | #587 | 3854816872 | Implemented: corrected MTP output contract docs | | #596 | 3846110059 | Already fixed on base: unambiguous GQA bias comment | | #600 | 3855174281 | Implemented: metadata-count-only MTP error | | #607 | 3855776048 | Implemented: stable route-field assertions | | #607 | 3855776086 | Implemented: public tensor iterator | | #609 | 3856486136 | Implemented: fail-closed LM-head comment | ## Current unresolved-thread disposition This covers all 48 Copilot threads returned by the reproducible #600-#630 query. The one human #623 thread is excluded. | PR | Comment | Current-main disposition and evidence | |---|---:|---| | #600 | 3855174177 | Already fixed by #629: package cycle and reserved-sidecar validation | | #600 | 3855174238 | Already fixed by #629: explicit MTP sidecar naming/loading | | #600 | 3855174281 | Implemented here: error no longer invents an observed block count | | #602 | 3848155961 | Outside exact stack; already fixed: top-level `expert_dtype` is classified before early return | | #602 | 3848155983 | Outside exact stack; still-valid behavioral block-quant validation, unchanged | | #602 | 3848156000 | Outside exact stack; still-valid truncated-read behavioral finding, unchanged | | #602 | 3848156022 | Outside exact stack; still-valid descriptor byte/dtype validation, unchanged | | #602 | 3848156040 | Outside exact stack; still-valid expert-bank payload validation, unchanged | | #603 | 3855249765 | Already fixed by #630: runtime preflight preserves shard sets | | #603 | 3855249840 | Already fixed by #630: success output follows durable runtime publication | | #604 | 3855343082 | Implemented here: graph-only MTP persistence distinguished from runtime rejection | | #604 | 3855343131 | Already fixed by #630: runtime success messages are atomic | | #607 | 3855776001 | Implemented here: missing generation golden skips before provenance read | | #607 | 3855776048 | Implemented here: only stable route fields are asserted | | #607 | 3855776086 | Implemented here: tensor count uses `tensor_items_raw()` | | #608 | 3856079371 | Still-valid behavioral cache-symlink containment finding; unchanged | | #608 | 3856079415 | Still-valid behavioral lowercase-digest validation finding; unchanged | | #609 | 3856486136 | Implemented here: comment matches value-preserving policy | | #610 | 3855541683 | Already fixed by #628: Falcon bias precedence is explicit | | #610 | 3855541761 | Already fixed by #628: CTRL tiny config exercises projection biases | | #611 | 3856595840 | Implemented here: runtime test resolves the distribution providing the module | | #612 | 3855677383 | Already fixed by #625: supported-version endianness detection | | #612 | 3855677427 | Implemented here: shared `INT64_MAX` sentinel | | #612 | 3855677460 | Implemented here: shared PLaMo2 width inference | | #612 | 3855677486 | Implemented here: accepted PLaMo2 activation spellings are explicit | | #613 | 3855845678 | Implemented here: canonical issue URL | | #613 | 3855845757 | Implemented here: Mamba-1 function-registration docs | | #613 | 3855845806 | Implemented here: test expects the canonical issue URL | | #614 | 3855988545 | Implemented here: removed stale Nemotron-H divergence comments | | #614 | 3855988597 | Already fixed by #628: zero-head geometry raises actionable `ValueError` | | #615 | 3856082290 | Already fixed by #626: dense GraniteHybrid bias closure | | #618 | 3856729330 | Implemented here: required routes filter ORT GenAI evidence | | #618 | 3856729409 | Implemented here: env-selected runtime version is authoritative | | #618 | 3856729490 | Stale/N/A: PR-description-only matrix claim; repository workflow claims one pinned version | | #618 | 3856729563 | Implemented here: schema tail restored to normal indentation | | #619 | 3856342484 | Already fixed on base: Kimi Linear uses `/issues/605` | | #619 | 3856342532 | Already fixed by #628: config rejects convolution kernels below 2 | | #619 | 3856342580 | Already fixed by #628: GGUF contract rejects convolution kernels below 2 | | #620 | 3855717041 | Already fixed by #627: tied LM-head-only checkpoints are retained | | #621 | 3856722324 | Already fixed by #628: Kimi-K3 required metadata is complete | | #623 | 3855931210 | N/A to current main: comment belongs to open, unmerged #623 | | #623 | 3855939302 | N/A to current main: comment belongs to open, unmerged #623 | | #623 | 3855939358 | N/A to current main: comment belongs to open, unmerged #623 | | #623 | 3855939394 | N/A to current main: comment belongs to open, unmerged #623 | | #624 | 3856777152 | Implemented here: runtime compatibility reuses the emitted model type | | #625 | 3857049733 | Implemented here: unsupported header reports both endian candidates | | #629 | 3857313182 | Newer post-audit behavioral sidecar-symlink cleanup finding; unchanged | | #629 | 3857313251 | Newer post-audit cross-platform path-safety finding; unchanged | ## Validation - affected GGUF/package/ORT GenAI/model/schema tests: 1,040 passed - broad non-integration suite: 7,851 passed, 56 skipped, 1 subtest passed - generated GGUF docs checks: 7 passed - initialized `lintrunner`; full lint/format passed - GPT-5.6 Sol medium review: one integrity finding fixed; re-review found no significant issues Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com> Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Summary
model.type: "decoder"contract for compatible single-model decoder-only text graphsstate_groupsValidation
onnxruntime-genai0.14.1 and 0.15.2Stacked on #609.